Papers with language decoder
Writing by Memorizing: Hierarchical Retrieval-based Medical Report Generation (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for medical image analysis use predefined template databases or ignore hierarchical nature of medical report generation. |
| Approach: | They propose a hierarchical retrieval mechanism to extract both report and sentence-level templates for clinically accurate report generation. |
| Outcome: | The proposed model extracts both report and sentence-level templates for clinically accurate report generation. |
LLMs Can Compensate for Deficiencies in Visual Representations (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a strong language backbone in vision-language models compensates for weak visual features by contextualizing or enriching them. |
| Approach: | They investigate whether strong language backbone compensates for weak visual features . they use CLIP-based vision encoders to perform controlled self-attention ablations . |
| Outcome: | The proposed model compensates for weak visual features by contextualizing or enriching them. |
Can Small Vision–Language Models Perform Sign Language Translation? (2026.findings-acl)
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have shown strong generalization across multimodal tasks, but their capacity to handle sign language translation (SLT) remains unclear. |
| Approach: | They propose entity- and semantics-aware metrics tailored for SLT to evaluate their performance. |
| Outcome: | The proposed metrics highlight the limitations of general-purpose VLMs to SLT, unlike their applicability in other tasks. |